Back

Clinical Trials

SAGE Publications

Preprints posted in the last 7 days, ranked by how well they match Clinical Trials's content profile, based on 11 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.

1
A Measurement-Based Care Strategy for Buprenorphine-Naloxone Treatment (Bup-MBC): Development of an EHR-Integrated Intervention

Reese, T.; Audet, C.; Ancker, J.; Wright, A.; Marcovitz, D.; Kast, K. A.; Bridges, J.; Tindle, H.; Shah, M.; von Horn, A.; Matheny, M. E.

2026-09-01 addiction medicine 10.64898/2026.08.27.26361539 medRxiv
Top 0.1%
18.6%
Show abstract

Introduction: Risk of recurrent opioid use during buprenorphine-naloxone (bup-nx) treatment is dynamic and remains elevated after initiation, with vulnerability shaped in part by treatment intensity and gaps between visits, yet routine outpatient care relies on episodic encounters and retrospective data. This mismatch can delay recognition of emerging instability and limit timely treatment adjustments. This paper reports the development and specification of an intervention strategy to address this mismatch. Methods: We used a structured, multi-phase design process to specify and configure a measurement-based care (MBC) strategy for bup-nx treatment (Bup-MBC) in outpatient addiction clinics through three phases: (1) a systematic review of patient-reported outcome measures (PROMs) for substance use treatment; (2) a qualitative needs assessment using the Theoretical Domains Framework and COM-B (Capability, Opportunity, Motivation-Behavior) model to identify gaps in risk monitoring, agency, and trust; and (3) iterative co-design with multidisciplinary clinicians to refine workflow fit and trust-preserving use of data. Patients informed item and feedback content during the needs assessment but did not participate in the co-design cycles. Results: Bup-MBC integrates (1) brief between-visit PROMs (e.g., withdrawal, craving, adherence); (2) immediate non-punitive patient feedback; (3) clinician-facing summaries and non-directive prompts in the electronic health record (EHR); and (4) an opt-in between-visit outreach pathway with predefined safety triggers, all configured within existing EHR and patient portal infrastructure. It targets patient and clinician capability to recognize changes in risk, opportunity for action through structured monitoring and visit preparation, and trust and agency through non-punitive communication, without adding substantial burden. The full measure set, severity bands, and question-to-action map are provided as supplementary material. Key trade-offs included prioritizing single-item measures for feasibility, balancing opt-in outreach with safety overrides, and assuming routine clinician use of summaries. Conclusion: This development study specifies an EHR-integrated MBC strategy for outpatient bup-nx treatment. As single-center design work with co-design limited to clinicians and delivery contingent on portal or text-message access, its outputs are hypotheses about mechanism and fit rather than demonstrated effects. Feasibility studies are needed to evaluate uptake, acceptability, workflow fit, and effects on treatment.

2
New tests for trials of very few patients using longitudinal data - a case-study in Autosomal Recessive Cerebellar Ataxias

Hendrickx, N.; Mentre, F.; Karlsson, M. O.; Hooker, A. C.; Traschütz, A.; Schüle, R.; PROSPAX Consortium, ; EVIDENCE-RND Consortium, ; Synofzik, M.; Comets, E.

2026-09-02 health informatics 10.64898/2026.08.28.26361588 medRxiv
Top 0.1%
6.6%
Show abstract

We propose two new tests to detect drug effects (DE) in trials of one to very few patients followed during two periods (before and after initiation of a treatment). Both methods use longitudinal natural history data to inform the estimation of each patient's DE. The first method uses a non linear mixed effect model (NLMEM) reflecting an expected natural history with a hypothetical drug effect, to estimate the Conditional Distribution of the Drug Effect (CDDE). The second method trains a Pareto Depth Analysis (PDA) algorithm, a machine learning based approach based on outlier detection, that we implement using data simulated under the NLMEM. We evaluated the two tests with a simulation study. We used data from the PROSPAX study in Autosomal Recessive Cerebellar Ataxias (ARCAs, to derive a NLMEM for the Scale for the Assessment and Rating of Ataxia score. The CDDE method provided controlled type I error and, in some scenarios, adequate corrected power, though sensitivity analyses showed vulnerability to misspecification. The PDA method demonstrated lower statistical power except with high score precision. These results highlight different strategies for quantifying treatment effects in ultra rare, patient' specific trials. They can inform methodological design for future ARCA precision therapies.

3
Bayesian Borrowing of External Information in Clinical Trials: A Comparison of MAP, RMAP, and SAM Priors

Choi, L.; McNeer, E.; Beck, C. A.; Neul, J. L.

2026-08-31 pharmacology and therapeutics 10.64898/2026.08.26.26360843 medRxiv
Top 0.1%
3.3%
Show abstract

Bayesian borrowing of external information can improve trial efficiency, particularly in pediatric and rare disease settings where patient populations are limited, but may introduce bias and inflate the Type~I error rate when the trial differs from external studies. Recent U.S. Food and Drug Administration (FDA) draft Bayesian guidance emphasizes careful evaluation of external information, prior specification, and assessment of operating characteristics. This paper compares three meta-analytic-predictive (MAP)-based methods for Bayesian borrowing: the MAP prior, robust MAP (RMAP) prior, and self-adapting mixture (SAM) prior. An adaptive platform trial design in Rett syndrome is used as a case study. Simulation studies evaluate frequentist operating characteristics under varying prior--data conflict, between-study heterogeneity, treatment effects, and clinically significant differences (CSDs) for the SAM prior. The MAP prior achieved the greatest efficiency when external and current data were compatible but exhibited the largest bias under substantial prior--data conflict. The RMAP priors improved robustness through fixed robust-component weights, whereas the SAM prior adaptively adjusted borrowing and was less sensitive to prior--data conflict while retaining efficiency gains when the data were compatible. Although the CSD influenced the degree of adaptive borrowing, as reflected by effective sample size, it had only a modest impact on frequentist operating characteristics. Sensitivity analyses using a skeptical robust component yielded similar qualitative conclusions, while accentuating the differences between the MAP and RMAP priors. These findings provide guidance for evaluating and selecting MAP-based borrowing strategies before trial implementation, particularly in rare disease settings, consistent with current FDA recommendations.

4
Increasing Lung Cancer Screening Participation Using an Informational Video Nudge: A Randomized Feasibility Trial

Wain, K. F.; Carroll, N. M.; Maclennan, A. J.; Hixon, B.; Steiner, J.; Ritzwoller, D. P.

2026-09-01 health systems and quality improvement 10.64898/2026.08.28.26361654 medRxiv
Top 0.1%
2.0%
Show abstract

Purpose: Lung cancer screening (LCS) with low-dose computed tomography (LDCT) reduces lung cancer mortality, yet screening participation remains low. We evaluated whether a brief informational video nudge delivered immediately before a scheduled clinical encounter increased LCS ordering and baseline LCS completion. Patients and Methods: We conducted a randomized feasibility trial within Kaiser Permanente Colorado from March through October 2025. LCS-eligible patients with an upcoming primary care or pulmonology appointment were assigned to intervention or usual care based on birth month. Intervention patients were split into two group, a group who received the LCS informational video nudge via text message within 24 hours of an eligible appointment; and second group who received the text plus a QR code video link during appointment rooming. Outcomes included LCS orders, baseline LCS-LDCT completion, and video engagement. Multivariable logistic regression was used to evaluate factors associated with LCS ordering. Results: Among 1,093 patients, 549 were assigned to intervention and 544 to usual care. Intervention patients were more likely to receive an LCS order within 1 day of their appointment (22.6% vs 16.4%; p=.010) and any time during follow-up (32.6% vs 24.1%; p=.002). Baseline LCS-LDCT completion was 51% higher in the intervention group, although the difference was not statistically significant (8.6% vs 5.7%; p=.078). Among the intervention group, 93 individuals (17%) viewed the video, generating 114 total views, and viewers watched an average of 79% of the video. Most views (82.5%) occurred through text-message delivery rather than QR codes. Conclusion: A brief, low-burden LCS informational video delivered immediately before a clinical encounter and integrated into existing workflows significantly increased LCS ordering and was associated with higher screening completion. Timely, scalable digital nudges may provide an effective strategy for improving LCS participation. Based on the observed effectiveness, feasibility, and efficiency of the intervention, KPCO incorporated the behavioral nudge into standard clinical care in February 2026.

5
Prospective In-silico Simulation of the VESALIUS-CV Trial Using Biomedical Knowledge Graph and Real-World Data-Driven AI Modeling

Perlman, A.; Goldstein, N.; Goldman, M.; Shapiro, M.; Barash, E.; Bar, A.; Raveh, T.; Tordjman, E.; Schussheim, H.; Dormont, F.; Matalon, O.

2026-08-31 cardiovascular medicine 10.64898/2026.08.26.26361436 medRxiv
Top 0.2%
1.3%
Show abstract

Background. Cardiovascular-outcomes trials are lengthy, costly, and associated with substantial uncertainty prior to readout. In-silico trial simulation using real-world data (RWD) has emerged as a potential tool to support earlier decision-making; however, evidence of prospective predictive validity, generated prior to trial result disclosure, remains limited. Methods. We applied a semi-mechanistic machine learning framework integrating real-world patient data with biologically informed drug representations to prospectively simulate the VESALIUS-CV trial evaluating evolocumab versus placebo. The simulation model was trained on a combination of patient-level real-world data and a drug-centric knowledge graph and validated for both patient-level and trial-level retrospective predictive performance. The model was then used to simulate VESALIUS-CV before public disclosure of trial results, using a locked model and prespecified eligibility criteria and primary endpoint aligned with the clinical protocol. A patient-level time-to-event model was used to generate virtual trial arms, from which cumulative incidence curves, hazard ratios, confidence intervals, and p-values for major adverse cardiovascular events (MACE) were estimated. Results. In retrospective validation, the model demonstrated strong patient-level discrimination, with time-dependent ROC-AUC values ranging from 0.80 to 0.90 across follow-up horizons. For trial-level validation, 22 randomized cardiovascular-outcomes trials were simulated, and hazard ratios for 3-point MACE across 24 between-arm comparisons showed consistent directional agreement and quantitative correlation with published results such that the model accurately predicted trial success, achieving an F1 score of 0.83, with precision of 0.79 and sensitivity of 0.89. In a fully prospective application, the simulation predicted a statistically significant reduction in 3-point MACE with evolocumab versus placebo, estimating a hazard ratio of 0.78 (95% CI, 0.70-0.87) at 54 months. These predictions were consistent with the subsequently reported VESALIUS-CV results, which demonstrated a hazard ratio of 0.75 (95% CI, 0.65-0.86) at 55 months of median follow-up. Conclusions. In a fully prospective setting, a RWD-driven, AI-based simulation accurately predicted the direction, magnitude, and temporal dynamics of treatment effects observed in the VESALIUS-CV trial. These results demonstrate that in-silico trial simulation can anticipate clinical outcomes in the prospective setting, supporting its use as a complementary tool for early decision-making, trial design optimization, and de-risking in cardiovascular drug development.

6
Integrated MB-PhD training is a long-term investment in the clinician-scientist workforce

Jafree, D. J.; Sun, M.; Stewart, G. W.; Gishen, F.; Swanton, C.; Motallebzadeh, R.; UCL MB-PhD Outcomes Study Group,

2026-08-31 health policy 10.64898/2026.08.26.26361003 medRxiv
Top 0.2%
1.1%
Show abstract

Background: Clinician-scientists translate clinical observation into discovery, trials, and policy, yet this workforce is shrinking across health systems worldwide. Integrated MB-PhD training, pausing medical training to complete a PhD before clinical exposure or specialisation, is one route into this career. We aimed to evaluate the long-term value of MB-PhD training and the barriers to clinical-academic careers these face after graduation. Methods: We evaluated all 131 graduates (29.8% female) who entered the University College London (UCL) MB-PhD programme over a 25-year period (1994-2018). Bibliometric outputs were collated via an inter-linked information system. Concurrently, all 131 graduates were invited to respond to open-ended questions on career benefits and structural barriers; 99 (75.6%) responded, and responses were independently coded into themes, which were then reviewed and confirmed by a Study Group of 107 individuals, including the 91 respondents who agreed to participate further. Results: Graduates produced 5,877 publications (1,141 first-author, 819 corresponding-author), attracting 350,754 citations, with a mean relative citation ratio of 3.30 {+/-} 0.47, approximately three times the field average and sustained across three decades of programme entry. Graduates secured an estimated $157.55 million across 99 grants, released 465 public datasets, and were named investigators on 31 clinical trials across five continents. Among the 99 survey respondents, 49.5% held consultant-grade posts, 72.7% remained research-active, and 25.3% had reached senior academic grade. Open-ended responses were coded into five recurring structural barriers, subsequently confirmed by the Study Group: insufficient protected research time (72.2% of responses), unsupportive training structures and limited career opportunities (36.7%, 24.4% of responses), funding and pay barriers (22.2% of responses), and lack of mentorship or geographical/family constraints (14.4%, 13.3% of responses). Conclusions: Integrated MB-PhD training generates sustained academic productivity and leadership, but structural barriers threaten retention of graduates within clinical-academic careers. Protecting research time, stabilising funding and pay, and reducing geographic instability are needed to retain the clinician-scientists that health systems have already invested in training.

7
Artificial Scientific Intelligence for Measurement-burden-aware Modelling and Interpretation of Multi-site Bone Mineral Density

Xiang, S.; He, H.; Xie, Z.; Cheng, C.-Y.; Li, H.; Liu, D.

2026-09-01 health informatics 10.64898/2026.08.30.26361665 medRxiv
Top 0.2%
1.0%
Show abstract

Agentic workflows can coordinate modelling, but balancing predictive performance, measurement burden and reproducibility is unclear. We developed DXA Agent, an agentic workflow for dual-energy X-ray absorptiometry (DXA) outcomes integrating planning, feature-model refinement, tools, provenance and hypothesis-generating interpretation. Models were independently developed and tested in UK Biobank (5,318 participants) and the National Health and Nutrition Examination Survey (NHANES; 3,777 participants), using cost-efficient and no-limit strategies. Across 20 UK Biobank and three NHANES bone mineral density sites, cost-efficient models achieved lower RMSE and higher R2 than the best conventional comparator, with median relative RMSE reductions of 10.9% and 9.9%, respectively. Classification was task dependent: UK Biobank osteoporosis averaged AUROC 0.839 and PR-AUC 0.182, whereas NHANES performance was comparable with conventional models. Higher-burden features did not consistently improve prediction. These retrospective, cohort-internal findings position DXA Agent as an inspectable, measurement-burden-aware research workflow requiring independent prospective validation.

8
Immune Checkpoint Blockade Modifies Drug-Associated Toxicity Across Phenotypes and Time

Mukherjee, E. M.; Asiaee, A.; Park, D.; Krantz, M. S.; Stone, C. A.; Martin-Pozo, M.; Phillips, E. J.

2026-09-02 dermatology 10.64898/2026.08.31.26361880 medRxiv
Top 0.2%
1.0%
Show abstract

Importance: Immune checkpoint inhibitors (ICIs) produce diverse immune toxicities, but whether checkpoint blockade also modifies associations between other drugs and adverse events is poorly understood. Objective: To define ICI-associated toxicity organization and determine whether drug-associated adverse events and onset vary with ICI exposure and checkpoint pathway. Design and Setting: Cross-sectional analysis of deduplicated FAERS reports from 2016 through 2025; analyses performed in 2026. Participants: Among 13,701,106 deduplicated reports, 2,365,269 were cancer associated and 256,940 contained an ICI. Median age among cancer reports with observed age was 66 years (IQR, 56-75 years); 1,031,999 (43.6%) were female and 1,003,154 (42.4%) were male. Exposures: ICI exposure in any reported drug role, individual primary-suspect drugs, and checkpoint-pathway exposure. Main Outcomes and Measures: Reporting odds ratios (ORs), cross-organ adverse-event communities, adjusted primary-suspect drug x ICI interaction ORs for Stevens-Johnson syndrome/toxic epidermal necrolysis (SJS/TEN), drug reaction with eosinophilia and systemic symptoms (DRESS), acute generalized exanthematous pustulosis (AGEP), interstitial nephritis, drug-induced liver injury (DILI), and vomiting (VOM), and accelerated failure-time model time ratios for documented onset. Results: Of 3001 eligible Preferred Terms in cancer-associated reports, 2091 differed at a false discovery rate (FDR) less than .05. Four cross-organ toxicity communities were identified. Of 138 eligible drug-phenotype pairs, 65 had FDR-significant interactions, including moxifloxacin-SJS/TEN amplification (interaction OR, 101.72; 95% CI, 39.11-264.55), enfortumab vedotin-SJS/TEN attenuation (interaction OR, 0.17; 95% CI, 0.13-0.23), and omeprazole-interstitial nephritis amplification (interaction OR, 10.35; 95% CI, 7.62-14.05). Among 60,324 reports contributing to temporal analyses, ICI exposure was associated with longer adjusted documented time to onset for 5 of 6 phenotypes (time ratios, 1.37-1.59) but not AGEP (time ratio, 0.99; 95% CI, 0.67-1.46). Temporal associations also differed across checkpoint pathways. Conclusions and Relevance: ICIs were associated with a structured cross-organ toxicity landscape, phenotype-specific modification of drug-associated adverse events, and distinct temporal patterns across checkpoint pathways. These findings support checkpoint blockade as a modifier of drug-associated toxicity and motivate longitudinal and mechanistic validation.

9
LLM-assisted evidence audit of late-stage cancer incidence as a screening trial endpoint

Li, S.; Zhang, W.; Xing, X.; Shen, Z.; Wang, Y.; Chen, Z.; Neto, O.; Yu, Y.; Wu, C.; Lin, L.

2026-08-31 oncology 10.64898/2026.08.29.26361733 medRxiv
Top 0.2%
0.9%
Show abstract

Background Late-stage cancer incidence is being considered as an earlier endpoint in cancer-screening trials, but its trial-level association with cancer-specific mortality may depend on evidence selection and endpoint harmonization. We evaluated the robustness of this association to source-verified additions. Methods We reconstructed the PubMed corpus underlying a 41-comparison review. Gemini 3.1 Pro Preview was used only to prioritize reports for blinded human reassessment. Reviewers determined eligibility, linked reports from the same trial, harmonized endpoints, and verified comparison-level data. We recalculated unweighted Pearson correlations overall and by cancer type after adding earliest-compatible trial comparisons. Results Among 1209 candidate records, 996 PDFs were assessed. Thirty-three reports absent from the source review were prioritized; 26 were eligible, representing 18 trials, and 8 provided compatible comparisons. Adding these comparisons increased the dataset from 41 to 49 and attenuated the overall correlation from 0.73 (95% confidence interval [CI] = 0.55 to 0.85) to 0.59 (95% CI = 0.37 to 0.75). Updated correlations were 0.49 (95% CI = -0.26 to 0.87) for breast, -0.23 (95% CI = -0.71 to 0.40) for colorectal, and 0.83 (95% CI = 0.54 to 0.95) for lung cancer. One sparse-event comparison influenced the colorectal estimate. Conclusions The overall association was sensitive to evidence composition, and cancer-specific stability varied. Late-stage incidence should be evaluated by cancer type and with prespecified sensitivity analyses for evidence selection and endpoint definitions. Model-assisted prioritization cannot replace human eligibility review, trial reconciliation, and source verification.

10
Half of alcohol, drug, and self-harm presentations cannot be identified in coded emergency department data: a diagnostic accuracy study of a large language model

Humphries, C.; Brett, J.; Gruber, F.; James, E.; McKendrick, T. I.; McNairn, K. C.; Miell, A.; O'Brien, R.; Rahman, F.; Schölin, L.; Stewart, M.; Casey, A.

2026-08-31 health informatics 10.64898/2026.08.26.26361443 medRxiv
Top 0.3%
0.9%
Show abstract

Objective To measure the accuracy of clinical coding, clinician review, and a locally deployed large language model (LLM) in identifying alcohol, drug, and self-harm involvement in emergency department (ED) attendances, and quantify prevalence. Design Two-phase diagnostic accuracy study. In a validation week, the identification strategies were assessed against a conflict-adjudicated reference standard (n=2,256); the LLM was then applied to n=105,096 annual attendances at the same site. Setting UK Type 1 Emergency Department treating patients [&ge;]16yrs. Main outcome measures Prevalence quantification compared with the reference standard; sensitivity, specificity, and balanced accuracy of each strategy; monthly identification rates and adjusted annual prevalence. Results The reference standard identified 12.1% of attendances as involving alcohol, drugs, or self-harm (coding 6.0%; clinician 10.0%, LLM 15.6%). LLM balanced accuracy matched or outperformed clinician review in all three domains (alcohol 0.942 v 0.930, p=0.635; drug 0.959 v 0.791, p<0.001; self-harm 0.982 v 0.908, p=0.004). Coding recorded 1.07 domains per identified patient against 1.32 in the reference standard. Adjusted annual prevalence corresponded to 12,890 domain involvements per year not identifiable in coded data. Subdomain classification found at least 81.6% of self-harm attendances required medical assessment for injury or overdose before psychiatric review. Conclusions Clinical coding identified fewer than half of presentations involving alcohol, drugs, and self-harm and rarely captured co-occurring domains; under-recording was present across a full year. A locally deployed LLM generated more complete structured data from existing clinical text within NHS infrastructure, at a scale which is not feasible for manual review.

11
Default-filled outcome labels in a deployed cognitive-screening programme: an operator-level audit and the construction of twenty-four language-model arms

Ji, J.; Sun, Z.; Ying, X.; Hao, J.; Fu, Z.; Shi, D.; Kong, X.; Xu, Y.; Zhang, X.; Du, X.; Zhang, Z.; Liu, X.; Lin, P.; Wang, H.

2026-09-02 health informatics 10.64898/2026.08.28.26361585 medRxiv
Top 0.3%
0.8%
Show abstract

Background. Routine service databases are attractive sources of training labels for clinical prediction models, but the processes that write those labels are rarely audited before the labels are used. In a deployed community cognitive-screening programme, we audited the routine cognitive-status label, built a matrix of twenty-four model arms over the same patients under a specialist reference standard, and measured what each supervision choice bought or cost. Methods. The study cohort is the 672 individuals whose cognitive status was recorded by a titled (attending-or-above) physician, that record being the reference standard; after holding out one institution entirely, a development panel of 642 individuals at 38 institutions. The routine cognitive-status label these individuals also carry was first audited at the operator level: for each data-entry account we counted diagnoses entered and the proportion recording any impairment, and tested a competing bulk-timestamp explanation. Twenty-four arms span the supervision choices such a programme faces: an incumbent 21-variable logistic regression; local language models (Qwen2.5-1.5B/3B, Qwen3-4B/8B) zero-shot, with chain-of-thought, fine-tuned on physician labels, on routine labels with and without decontamination, or on a proxy scale-band task; preference-optimised (DPO) and reinforcement-trained (GRPO) variants; a proprietary frontier model queried zero-shot; and knowledge distillation of that frontier model into the regression and into the local 4B, using 943 teacher-labelled records from the programme's unlabelled pool. All arms are scored out-of-fold under one five-fold split grouped on registry-resolved institution clusters (no cluster spans a fold); paired contrasts use a 2,000-draw cluster bootstrap. Results. 181 operator accounts (each entering at least 100 diagnoses with zero recorded impairments) account for 45,315 rows - 40.5% of the outcome column; recorded impairment falls monotonically with account volume (15.7% for 1-9 rows to 0.7% for 500-999); a bulk-timestamp explanation was tested and refuted, identifying the write-time column as a migration artefact. Under the specialist standard, no locally fine-tuned arm beat the incumbent regression (AUROC 0.926): physician-label SFT reached 0.924 (4B), DPO 0.881, and GRPO 0.789; the pre-registered two-stage proxy-then-RL recipe was worse than its single-stage contaminated baseline (-0.030, 95% CI -0.077 to -0.004). Chain-of-thought reduced discrimination at every size (-0.072, -0.080, -0.041 at 1.5B/3B/4B; -0.012, n.s., at 8B). The frontier model scored 0.932 (vs. regression +0.007, n.s.). The distilled 4B reached 0.940 - above the incumbent (+0.014, 0.004 to 0.031) and above its own teacher (+0.008, 0.001 to 0.017) - with near-teacher calibration; it reached the teacher's level by 50 teacher labels and changed little beyond 200. Conclusions. The audit and the arm matrix support one deployment recipe: audit the routine label at the operator level before training on it; do not expect fine-tuning, preference optimisation, or reinforcement learning on a few hundred specialist cases to beat a well-calibrated regression; and if a frontier model is available but undeployable, spend a bounded number of queries on it as a labelling instrument and distil. A companion paper uses these frozen predictions to quantify how evaluation design choices compare with model choice.

12
Large language model-augmented implicit surgical video review

Zhang, Z.; Qadir, M. I.; Ramchand, R.; Belwadi, M.; Ball, R. P.; Konstantinopoulos, K.; Abbey, E. M.; Ernsberger, K. T.; Guzman, M. J.; Hendren, S.; Holcomb, B. K.; Robb, B. W.; Stankowski, T.; Waters, J. A.; Stefanidis, D.; Bilimoria, K. Y.; Mohanty, S.; Kolbinger, F. R.

2026-08-31 surgery 10.64898/2026.08.25.26361071 medRxiv
Top 0.4%
0.5%
Show abstract

Surgical video interpretation is a promising medical artificial intelligence application. However, no existing video annotation method preserves the spatiotemporal complexity of surgeon reasoning. Here we show that verbal reasoning and visual attention can be converted into structured, machine-actionable records of intraoperative behaviours. Our method decomposes transcribed verbal commentary into video-anchored semantic feedback chunks, which are classified via a large language model, with spatial grounding to surgical scenes via eyegaze or cursor tracking. We demonstrate method validity and scalability on structured and unstructured annotation tasks. For quality feedback on full-length colorectal procedures, the method reached near-human fidelity for chunking (mean cosine similarity: 0.95, SD: 0.01) and semantic classification across observations (mean Cohen's kappa: 0.71, SD: 0.07) and evaluative triggers (mean Cohen's kappa: 0.67, SD: 0.14), with excellent usability ratings. For structured critical view of safety assessment in laparoscopic cholecystectomy, implicit annotation yielded excellent agreement with explicit reviewer ratings (Cohen's kappa: 0.83, 0.49 and 0.81 across three criteria). We anticipate this method will advance surgical data science by enabling scalable construction of meaningfully annotated surgical video datasets.

13
When medical credentials conflict with stated accuracy: A factorial study of source credibility and answer revision in medical LLM interactions

Wojcik, S.; Rulkiewicz, A.; Domienik-Karłowicz, J.

2026-09-01 health informatics 10.64898/2026.08.28.26361634 medRxiv
Top 0.4%
0.5%
Show abstract

Large language models perform well on medical examinations, but users routinely challenge their answers and invoke professional roles, and it is unclear what a system does when a medical credential and a stated task-specific accuracy point in opposite directions. In a factorial experiment on 480 items from four Polish specialty examination sets and three consumer large language model systems (ChatGPT, Claude, Gemini), each item and system received eleven independent conversations. Conditions crossed attributed source role (medical student, experienced specialist), stated prior accuracy on similar questions (2/10, 8/10) and suggestion correctness. The primary outcome was adoption of a prespecified incorrect option when the baseline answer matched the official key, comparing a specialist described as 2/10 with a student described as 8/10. Baseline agreement with the key was 87.2% across 15,683 analyzable conversations. The incorrect option was adopted more often from the specialist described as 2/10 than from the student described as 8/10 (10.2% vs. 7.6%; adjusted risk difference +2.82 percentage points, 95% CI +0.65 to +4.99). Estimates varied across the three systems and only one system-specific interval excluded zero. In a prespecified exploratory analysis with a shared eligibility rule, correct suggestions were adopted far more often than incorrect ones (risk difference +35.7 percentage points, 95% CI +30.8 to +40.7), indicating selective rather than indiscriminate compliance. An incorrect suggestion from a specialist with low stated accuracy was therefore slightly more influential than the same suggestion from a student with high stated accuracy, although the difference was modest and varied across systems. Agreement reached only after a user has disclosed a preferred answer should not automatically be treated as an independent second opinion, and medical large language model systems should be evaluated on how they revise answers after such disclosure, not solely on initial accuracy.

14
The use of computerised testing to assess cognitive performance in people with HIV in South Africa

Edmond, E. C.; Dreyer, A. J.; Winston, A.; Khoo, S. H.; Joska, J.; Nightingale, S.

2026-08-31 hiv aids 10.64898/2026.08.27.26361083 medRxiv
Top 0.4%
0.5%
Show abstract

Background Computerised cognitive testing may address the global challenge in identifying cognitive changes in people living with HIV scalably and affordably. We assessed a computerised battery (CB) of cognitive tests, in a prospective cohort (CONNECT) of people with HIV in a low-income peri-urban area of Cape Town, South Africa during a national programmatic switch from efavirenz- to dolutegravir-based antiretroviral therapy (ART). Methods We recruited 170 people with HIV and 91 people without HIV (controls) (140[82%] and 41[45%] followed up). The CB and gold-standard pen&paper cognitive testing (P&P) were performed at both timepoints. Technology familiarity/use questionnaire data were also collected. We compared performance in detecting lower group-level cognitive performance associated with efavirenz treatment. Furthermore, the CB was compared to P&P in classifying individuals with low cognitive performance, correlation of global test scores and domain-level scores between batteries, and practice effects between timepoints. Exploratory principal component analysis was also performed. Results People with HIV on efavirenz at baseline had lower performance on the computerised battery than controls, {Delta}T=2.6, p=0.0047. This difference was lost after switching to dolutegravir-based ART at follow-up. CB and P&P global T were moderately correlated (R2=0.203, p<0.001), and the CB performed moderately in classification of low cognitive performance against the gold standard (AUC 0.70, sensitivity 0.52, specificity 0.76, PPV 0.40, and NPV 0.84). Selecting the first three principal components improved both classification of low cognitive performance (AUC 0.77) and correlation strength with P&P global T (R2=0.3, p<0.001). The CB did not show practice effects. Most participants owned a mobile phone (95%, 85.9% of these smartphones). Performance was better in smartphone owners ({Delta}T=1.8) and computer owners (23%, {Delta}T=1.8). Conclusions Delivering computerised cognitive testing was feasible in this low-income southern African setting. The CB showed reasonable construct validity (detecting known lower cognitive performance associated with efavirenz-ART) and may detect broad cognitive characteristics such as processing speed and accuracy. However, correlation of CB results with gold standard P&P testing was low-moderate and may limit its applicability as a diagnostic tool. This might be improved by including a wider range of cognitive domains tested in the CB, or data driven analysis. Brief CBs may fulfil an initial screening role, followed by more detailed clinical assessment.

15
Best Practice Manufacturing and Quality Standards for Bacteriophage Therapy Products: Australian Consensus Statements

Watts, K.; Lin, R. C.; Lynch, S.; Warning, J.; Barr, J. J.; Ben Zakour, N.; Campbell, A.; Chan, J.; Collie, L.; Hedges, M.; Hudson, B.; Irwin, A.; Khatami, A.; Kicic, A.; Laucirica, D.; Lauter, C.; Ling, K.-m.; Ng, R.; Pavuk, N.; Rahmatullah, R.; Sinclair, H.; Tucker, E.; Vreugde, S.; Warner, M.; Velickovic, Z.; iredell, j.

2026-08-31 public and global health 10.64898/2026.08.26.26361487 medRxiv
Top 0.4%
0.5%
Show abstract

Objective As antimicrobial resistance (AMR) continues to threaten global public health, bacteriophage therapy products (BTPs) offer a promising alternative to conventional antimicrobials. However, translation into routine clinical practice requires best practice standards for manufacturing and quality control to ensure the consistent safety, quality, and reliability of personalised BTPs produced for individual patients or small cohorts. Design A modified Delphi methodology was used to develop consensus statements, engaging experts from Australia's National Bacteriophage Therapy Regulatory Working Group across the fields of clinical microbiology, phage biology, good manufacturing practice (GMP), regulatory science, and government. The process comprised three iterative phases: (1) structured statement development, (2) an anonymous REDCap survey, and (3) a hybrid consensus meeting. The strength of evidence and recommendations was assessed using the GRADE (Grading of Recommendations Assessment, Development and Evaluation) framework. Results Consensus was reached on 35 statements to provide best practice manufacture and quality control guidance for BTPs. These statements address requirements for phage identification and characterisation; define the point at which GMP-aligned processes commence for ubiquitous phages; outline quality control expectations for phage active pharmaceutical ingredient (pAPI) production and maintenance of BTP and host cell repositories. Additional guidance covers quality management systems, including documentation, traceability, and governance. Conclusion These consensus statements provide comprehensive best practice recommendations for the manufacture and quality control of BTPs in Australia. By promoting consistent, safe, and quality-assured approaches to personalised BTPs, they aim to facilitate clinical implementation while remaining aligned with existing international pharmacopoeial standards and regulatory frameworks.

16
Nurture Early for Optimal Nutrition (NEON): A pilot cluster randomised controlled trial of community-facilitator-led participatory learning and action womens groups to improve infant feeding & care among South Asian families in East London

Manikam, L.; Fatima, A.; Patil, P.; Mayadewi, C. A.; El Khatib, T.; Drazdzewska, J.; Oyebode, O.; Llewellyn, C. H.; Webb-Martin, K.; Irish, C.; Archibong, M.; Gilmour, J.; Kalungi, P.; Batura, N.; Shringarpure, K.; Lakhanpaul, M.; Heys, M.; NEON Steering Team,

2026-08-31 public and global health 10.64898/2026.08.28.26361604 medRxiv
Top 0.5%
0.4%
Show abstract

South Asian communities in the UK experience disproportionate maternal and child health inequalities linked to non-recommended infant feeding practices, limited health literacy, and socioeconomic constraints. Participatory learning and action (PLA) is effective in low- and middle-income countries, but high-income evidence is scarce. This pilot assessed the feasibility of a community facilitator-led PLA intervention to improve infant feeding among South Asian families in East London. A three-arm pilot feasibility cluster randomised controlled trial (ISRCTN10234623) was conducted in Tower Hamlets and Newham, East London (May-September 2022), with 12 wards randomised 1:1:1 to face-to-face PLA, online PLA, or usual care. Multilingual community facilitators delivered eight biweekly sessions over 14 weeks. Feasibility outcomes were assessed against prespecified Go/Stop criteria; exploratory outcomes included child feeding behaviours (Children's Eating Behaviour Questionnaire, CEBQ), parental feeding style (Parental Feeding Style Questionnaire, PFSQ), and child BMI Z-scores. Of 263 enrolled participants, 261 had a recorded trial arm allocation; consent to the pilot feasibility study was 70.7% (186/263; 95% CI 65.0-75.9%) meeting the [&ge;]50% Go criterion. Attendance was 37% (Tower Hamlets 59%, Newham 29%), below the [&ge;]80% Go threshold. Six-month retention was 54.8% (Tower Hamlets 78%, Newham 48.5%; 95% CI 41.8-55.3%), triggering the Definite Stop criterion. Significant baseline imbalances included BMI Z-score (p = 0.005), ethnicity, borough, and education; no between-arm BMI differences were observed at follow-up (p = 0.249). CEBQ and PFSQ baseline completion was 24.5% and 23.0%, with no usable follow-up data. PLA Phases 3 and 4 were not completed by any group; all participants providing feedback reported it acceptable. Recruitment was feasible and the intervention acceptable, but a Definite Stop criterion was triggered in Newham, no group completed the full PLA cycle, and outcome data were insufficient for evaluation. A definitive trial requires stratified randomisation, digitised multilingual data collection, participant reimbursement, and explicit PLA phase-completion criteria.

17
CHARMS and PROBAST+AI: an updated template for Data Extraction and Risk of Bias Assessment in systematic reviews of prediction models

Jaber, A.; Hughes, L.; Cameron, A. C.; Quinn, T. J.

2026-08-31 cardiovascular medicine 10.64898/2026.08.26.26361189 medRxiv
Top 0.5%
0.4%
Show abstract

Background: Systematic reviews of clinical prediction models increasingly include studies using artificial intelligence (AI) and machine learning (ML) methods alongside traditional multivariable regression approaches. A previously published Excel tool enabled standardised data extraction using the CHARMS checklist and risk of bias assessment using PROBAST. The recent publication of the PROBAST+AI framework, which distinguishes the assessment of model development quality from the assessment of model evaluation risk of bias and assesses applicability in both parts, necessitates an updated digital instrument applicable across prediction modelling methods. Methods: We updated an open-access Excel tool to incorporate the full PROBAST+AI framework. The updated template incorporates structural separation between assessment of model development quality and model evaluation risk of bias, with applicability assessed in both parts. It also incorporates updated signalling questions, including those addressing methodological issues particularly relevant to AI/ML, and automates the generation of summary tables and graphical displays. Results: The updated tool (CHARMS & PROBAST+AI Template) contains 11 worksheets and supports data extraction and appraisal for up to 30 prediction models. Dedicated, linked worksheets enable separate assessment of model development and model evaluation, with Domain 4 distinguishing among Apparent, Internal, and External evaluation settings. Key updates include dedicated assessments for predictor pre-processing, class imbalance handling and recalibration, data leakage prevention, and replication of the full model development pipeline within resampling procedures. Automated sheets dynamically format tables and summary charts covering PROBAST+AI parts. Conclusions: The CHARMS & PROBAST+AI Excel template provides a standardised, user-friendly, and rigorous digital framework for systematic reviewers appraising traditional statistical and AI-driven clinical prediction models.

18
Effects of collaborative clinical visit agenda-setting interventions: A systematic review and meta-analysis

Sierpe, A.; Yen, R. W.; Milliman, A.; Cady, E.; Ahn, B.; Dade, A. E.; Devito, A. M.; Eckert, B. A.; Gopalan, V. V.; Krasinski, S. C.; MacMartin, M. A.; Musacchio, S. G.; Zhang, J.; Saunders, C. H.

2026-09-03 medical education 10.64898/2026.08.30.26361729 medRxiv
Top 0.5%
0.4%
Show abstract

Background Agenda-setting is a fundamental patient-centered communication practice in which a clinician works with a patient to elicit, propose, and organize topics for discussion during a clinical encounter. Various agenda-setting interventions have been developed, including patient-facing tools and clinician training, but their effects have not been systematically evaluated. We aimed to determine the effects of these interventions on encounter, patient, care partner, and clinician outcomes. Methods We searched grey literature and seven databases, including PubMed, from inception through July 2025 for randomized and non-randomized comparative studies of interventions designed to promote or improve clinical visit agenda-setting. Two reviewers independently screened articles and extracted data, with a third reviewer resolving conflicts. We assessed risk of bias using RoB 2 for randomized studies and ROBINS-I for non-randomized studies. We conducted random effects meta-analyses when outcomes were sufficiently comparable, assessed heterogeneity using I2, and rated certainty of evidence using GRADE. Post hoc exploratory subgroup analyses examined study design, adjustment status, and intervention structure. Results Twenty-nine articles describing 22 unique studies met the inclusion criteria, including 13 randomized and nine non-randomized studies. Agenda-setting interventions increased the occurrence of agenda-setting (risk ratio 5.43, 95% confidence interval (CI) 2.06 to 14.28, I2=34.6%) and favored the intervention for concerns addressed when measured as a continuous outcome (standardized mean difference (SMD) 0.37, 95% CI 0.16 to 0.57, I2=65.3%) and overall clinician satisfaction (SMD 0.50, 95% CI 0.23 to 0.78, I2=0.0%). There were no clear differences in the number of concerns raised (mean difference (MD) 0.21, 95% CI -0.19 to 0.61, I2=59.6%), visit duration (MD 0.64 minutes, 95% CI -0.83 to 2.12, I2=51.4%), or overall patient satisfaction (SMD 0.05, 95% CI -0.05 to 0.15, I2=47.0%). Potentially important heterogeneity was present for four of these six outcomes. Post hoc exploratory subgroup analyses did not provide clear evidence that effects varied by study design, adjustment status, or intervention structure. Risk of bias was often high, serious, or critical, and certainty of evidence was low or very low for all pooled outcomes. Conclusions To our knowledge, this is the first comprehensive synthesis of clinical visit agenda-setting interventions. Such interventions may increase the occurrence of agenda-setting and the extent to which patient concerns are addressed without increasing visit length. However, the certainty of evidence was low or very low, and the available evidence does not establish a superior intervention structure.

19
Can Dental AI Really Beat Dentists? DentalPair-Cert for Rigorous AI-Dentist Inference

Alve, S. R.; Rahman, S.; Meem, S. M. A. C.

2026-09-02 dentistry and oral medicine 10.64898/2026.09.01.26361874 medRxiv
Top 0.5%
0.4%
Show abstract

A dental AI system and a dentist reading the same radiographs form a paired comparison. Published comparative studies often report the two arms separately against a reference standard, leaving the joint pattern of correctness between them unavailable for secondary paired inference. We show what that omission costs. The accuracy difference remains exactly identified; its sampling variance does not, so the report contains the estimate and not its uncertainty. On a study of 282 units, two published accuracies are consistent with 38 distinct joint tables whose confidence intervals differ in width by a factor of 2.5. The consequence is a three-zone decision map rather than a single threshold: differences at or below 1.06 points are non-significant under every compatible table, differences at or above 6.03 points are significant under every compatible table, and in between the published numbers cannot decide. We then show the omission is repairable at negligible cost. One additional integer, the number of units both arms classify correctly, identifies the joint table exactly and restores standard paired inference. For a panel of readers the pairwise dependences must arise from one joint distribution, a constraint that binds once three readers are present; publishing each reader's joint-correct count against a single reference reader cannot widen and may tighten every pairwise bound, and in a 7-arm experiment reduced them by a median of 37% even for pairs excluding that reference. Where the integer was never published we give DentalPair-Cert, an interval with finite-sample coverage uniformly over every admissible within-unit AI-dentist dependence under the independent-sampling-unit model, certified in both the nuisance maximization and the inversion. Across 4,200,000 simulated comparisons an independence analysis falls to 74.5% coverage with 12.2% type-I error; in a purposive sample of 9 recent comparative studies, 1 reported a paired test on discordant units.

20
CPT/HCPCS Code Recommendation from Clinical Notes: A Comparative Evaluation of AI Methods

Song, Q.; Ni, C.; Liu, W.; Li, Y.; Malin, B. A.; Yin, Z.

2026-08-31 health informatics 10.64898/2026.08.29.26361731 medRxiv
Top 0.5%
0.4%
Show abstract

Automatic coding from clinical notes has been studied extensively for International Classification of Diseases (ICD) codes, yet broad Current Procedural Terminology (CPT) and Healthcare Common Procedure Coding System (HCPCS) recommendation remains comparatively underexplored. Existing studies often focus on one specialty, a limited code vocabulary, or a single model family, leaving it unclear how different artificial intelligence (AI) paradigms perform under a common, clinically meaningful evaluation. We formulate CPT and HCPCS coding as an AI-assisted recommendation task in which a physician or professional coder reviews a short, ranked list of candidate codes supported by the clinical note. Using operative notes from Vanderbilt University Medical Center (VUMC) and discharge summaries from Medical Information Mart for Intensive Care IV (MIMIC-IV), we compare lexical retrieval, Clinical-Longformer, GPT-5.6-Sol, MedGemma-27B, and an inspectable agentic-style retrieve-and-verify system under a controlled review budget. Micro-averaged recall within a fixed number of recommendations measures whether reference codes reach the reviewable list; micro-F1 is reported only where reference labels are sufficiently complete. Zero-shot GPT-5.6-Sol achieves the highest recall within five and ten candidates: 0.717 and 0.800 on VUMC and lower-bound values of 0.689 and 0.738 on MIMIC-IV. The retrieve-and-verify system reaches 0.695 and 0.784 on VUMC and lower-bound values of 0.575 and 0.657 on MIMIC-IV, with a candidate-linked evidence window attached to each retained recommendation. Diagnostic analyses reveal distinct failure sources, including output-length underfilling, confusion among closely related codes, out-of-knowledge-base generation, and incomplete evidence support. These findings establish a systematic evaluation framework for procedure-code recommendation and identify practical requirements for future systems that are accurate, review-efficient, and grounded in clinical evidence.